The Journal of the Acoustical Society of America
● Acoustical Society of America (ASA)
Preprints posted in the last 30 days, ranked by how well they match The Journal of the Acoustical Society of America's content profile, based on 35 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Bozdogan, A.; Aarts, R. M.
Show abstract
Elephants and other large mammals produce low-frequency vocalizations extending well below the 20 Hz lower limit of human hearing, a regime known as infrasound. These rumbles serve vital social and reproductive functions over distances of several kilometers, yet they are inaudible to human observers and cannot be reproduced by conventional small loudspeakers. We present a complete signal-processing pipeline that renders sub-20 Hz elephant rumbles perceptible through a small loudspeaker by exploiting the missing-fundamental psychoacoustic effect. Butterworth bandpass filters isolate the infrasonic content; a full-wave integrator nonlinear device (NLD) generates the harmonic series required for virtual pitch perception; and a hysteresis-comparator fundamental-frequency estimator normalizes the NLD output. The pipeline was validated on African elephant field recordings and deployed on a credit-card-sized, low-cost single-board computer with an infrasound microphone and a small Bluetooth loudspeaker, demonstrating live operation in the field. The processed output shows a 10 dB to 15 dB elevation in the loudspeakers efficient band during call segments compared with background. The system enables zoo visitors and wildlife observers to perceive elephant rumbles in real time, opening new avenues for behavioral studies and public engagement with animal communication.
Dirks, C. E.; Guest, D. R.; Oxenham, A.
Show abstract
Context effects are ubiquitous across sensory systems and reflect a general encoding principle for both simple and complex stimuli. One simple context effect, contraction bias, manifests in two-interval perception tasks as a bias of the perceived magnitude of the first stimulus toward the center of the overall magnitude range. The underlying cause of contraction bias is unclear. One explanation is that a listeners magnitude estimate of the first stimulus is combined with a perceptual anchor, usually the mean stimulus magnitude, biasing it toward the anchor (sensory model). An alternative explanation is that a listeners response criterion shifts, based on the magnitude of the stimulus pair, relative to the mean magnitude of the stimuli range (decision model). Two pitch-discrimination experiments were performed to test these hypotheses in the auditory domain. The first was a forced-choice discrimination task, where listeners were asked to identify the higher or lower tone in a pair. The second was a same-different task where listeners indicated whether or not the two tones in a pair differed in frequency. Contraction bias was observed in the higher-lower discrimination task, even after extensive perceptual training with feedback. In contrast, no contraction bias was observed in the same-different task. Computational models of the sensory and decision hypotheses were fit to data from both experiments. The sensory model captured the pattern of results the higher-lower experiment but erroneously predicted a contraction bias in the same-different task. The decision model produced similar predictions to the sensory model in the higher-lower task but correctly predicted no contraction bias in the same-different task, and produced lower prediction errors and more stable parameter estimates in both paradigms. Overall, the results suggest that the underlying nature of the contraction bias may reflect decision, rather than sensory, biases based on the context.
Marrone, J. P.; Ziliak, M. C.; Bartlett, E. L.
Show abstract
Auditory brainstem responses (ABRs) are a core part of objective functional evaluations of hearing sensitivity and subcortical auditory transmission. Manual assessments of ABR waveforms are still a primary means by which thresholds and peak amplitudes and latencies are measured, which is time-consuming and prone to user variability. Automated methods have offered promising alternatives for ABR classification, but they have sometimes been limited in accuracy or robustness. Here, we developed and tested a supervised convolutional neural network (CNN) based ABR peak classifier that works across sound levels and sound frequencies that can be run quickly on a personal computer using single or dual-channel ABR inputs. For ABR peaks I, III, IV, and V, the classifier achieved over 95% accuracy. High accuracy was maintained even after noise-exposure causing temporary or permanent threshold shifts, and over 90% of peaks were within 0.041 ms (1 sample) of the manually identified peak. Only a few hundred samples were needed to train the network, making it widely amenable to smaller data studies or where the number of subjects or sessions may be low.
Galeano-Otalvaro, J.-D.; Dieudonne, B.; Francart, T.; Wouters, J.
Show abstract
Understanding speech in noisy environments relies strongly on binaural cues such as interaural time differences (ITDs) and interaural level differences (ILDs), which support spatial hearing and the segregation of competing sound sources. When these cues are degraded, listeners experience substantial difficulty in complex acoustic environments. Behavioural measures of binaural benefit, such as binaural masking level differences (BMLDs), binaural intelligibility level differences (BILDs), and spatial release from masking (SRM), are well established in normal-hearing (NH) listeners, but they require an active behavioural response. Neural speech tracking using electroencephalography (EEG) has emerged as a promising approach for quantifying neural processing of continuous speech, yet its sensitivity to spatial hearing cues remains insufficiently characterised. In this study, we investigated the neural correlates of spatial release from masking in NH listeners using EEG-based neural speech tracking. Nineteen participants listened to continuous Dutch speech stories presented with masking noise under two spatial configurations, collocated (S0N0) and spatially separated (S0N90), across multiple signal-to-noise ratios (SNRs). Neural tracking of the speech envelope was quantified using both envelope reconstruction and temporal response function (TRF) analyses. Spatial separation enhanced neural tracking of the target speech envelope, particularly at challenging SNRs where behavioural SRM was also observed. TRF analysis further revealed condition-dependent morphologies, including increased amplitudes and decreased latencies of late cortical components consistent with spatial unmasking effects. These neural differences were most pronounced at low SNRs, where spatial cues provide the greatest perceptual benefit. Together, these findings demonstrate that neural speech tracking captures cortical signatures of spatial unmasking and closely reflects behavioural improvements in speech understanding. Establishing these relationships in NH listeners supports the development of objective neural measures for evaluating binaural benefit in difficult-to-test populations.
Fan, X.; Mathiassen, S. E.; Johansson, P. J.; Jackson, J. A.; Nyman, T.
Show abstract
This study examined how tempo, dynamics, and string influence upper-extremity physical exposure in professional violinists and how exposure variability is distributed among musical characteristics, between-subject differences, and residual variability. Twelve violinists performed seven standardized scales while bilateral upper-arm and wrist kinematics and shoulder and forearm muscle activity were recorded. Linear mixed-effects models showed that faster tempo increased right upper-arm velocity and bilateral forearm activity while reducing right upper-arm and wrist ranges of motion. Louder dynamics increased bilateral forearm and right trapezius activity and right-wrist ranges of motion. Higher-posture strings increased right upper-arm elevation and right shoulder muscle activity. Variance analysis identified exposures predominantly related to musical characteristics, jointly related to musical characteristics and between-subject differences, predominantly related to between-subject differences, or mainly unexplained. These findings support future exposure prediction from musical characteristics and targeted prevention through repertoire-based workload management, structured recovery, and individualized technique-focused strategies.
Azadpour, M.; Neukam, J.; Capach, N.; Svirsky, M.
Show abstract
Cochlear implants (CIs) restore hearing by stimulating auditory neurons to encode amplitude envelopes across frequency bands, providing essential cues for speech recognition. This study investigated how stimulation pulse rate constrains temporal envelope processing and speech cue perception in ten post-lingually deaf CI users by evaluating amplitude modulation (AM) detection thresholds and consonant identification performance across pulse rates. The effects of pulse rate on temporal processing and speech perception were examined using both standard clinical multi-channel strategies and single-channel strategies designed to isolate within-channel envelope representations. Results revealed a significant decline in AM detection and consonant recognition performance at the lowest tested pulse rate of 125 pulses per second (pps), consistent with perceptual constraints on temporal processing at low carrier rates, rather than inadequate envelope sampling. At the highest pulse rate of 4000pps, a non-significant reduction in AM detection was observed which may be consistent with previously reported reductions in amplitude discrimination at high pulse rates. Consonant recognition performance remained stable across clinically relevant pulse rates (250-2000pps), though listener-specific pulse rate effects were observed. Notably, significant correlations were found between single-channel and multi-channel performance in AM detection and consonant recognition tasks. These findings support an important contribution of within-electrode temporal envelope processing to multi-channel speech perception and highlight the clinical relevance of individual variability in pulse rate effects.
McCorkendale, B.; Rodriguez, R.; Fink, R.; Moore, M.; Romero, S.; Esmailie, F.
Show abstract
PurposeMild therapeutic hypothermia (MTH) preserves cochlear function in animal models and is now entering early-phase human trials for hearing preservation. However, the extent to which the human cochlea can actually be cooled, and the mechanisms underlying MTH, remain unclear, in part because blood perfusion is expected to oppose localized cooling. In this study we evaluated the impact of blood flow on human cochlear temperature exposed to the MTH device using a combined experimental and computational approach. MethodsTemperature measurements were obtained from a human cadaver skull exposed to a commercial MTH device. These data were used to validate a three-dimensional bioheat transfer model incorporating realistic skull anatomy. The validated model was subsequently extended to include physiological blood perfusion in the internal carotid artery; a major heat source located near the cochlea. Finally, the in silico model was further expanded to incorporate the surrounding skin and brain tissues. ResultsIncorporating blood flow in internal carotid artery substantially altered predicted cochlear temperature distributions, highlighting the importance of localized vascular heat transport in the human cochlea during MTH. Although cochlear cooling was attenuated in the presence of perfusion, the therapeutic effects of MTH may not depend solely on the magnitude of local intracochlear temperature reduction. Additional mechanisms, such as reduced facial surface temperature, may also contribute to its efficacy. ConclusionThe validated in silico model provides a physiologically realistic framework for evaluating human cochlear thermal responses, investigating MTH mechanisms, and optimizing temperature-based strategies for hearing preservation.
Wang, F.; Utianski, R. L.; Duffy, J. R.; Barnard, L. R.; Botha, H.
Show abstract
This study examined the extent to which goodness of pronunciation (GoP) scores and phonological posterior probabilities capture perceptual ratings of speech severity in individuals with motor speech disorders (MSD). Speech recordings of the word catastrophe were obtained from 489 participants, including 333 neurologically typical controls and 156 individuals with MSD. GoP scores were derived using traditional acoustic features and self-supervised speech representations, including WavLM and XLS-R, across multiple modeling approaches, while phonological posterior probabilities were extracted using Phonet. Model performance was evaluated using Kendall's rank correlations, regression, and receiver operating characteristic analyses against speech-language pathologists' perceptual ratings of sound distortion and intelligibility. Both GoP and phonological posterior probabilities were significantly associated with perceptual ratings. Self-supervised speech representations substantially outperformed traditional acoustic features, with WavLM-based GoP using k-nearest neighbors achieving the strongest performance. Across correlation, regression, and classification analyses, GoP consistently outperformed phonological posterior probabilities for both sound distortion and intelligibility. Age and gender had minimal influence on model-derived measures or their relationships with perceptual ratings. These findings demonstrate the value of self-supervised GoP as an objective measure of speech impairment while highlighting the complementary role of phonological posterior probabilities in characterizing articulatory aspects of motor speech disorders.
Nidiffer, A.; O'Sullivan, A.; Lalor, E. C.
Show abstract
In noisy environments, visible speech articulations improve listening comprehension. The benefit derives from several sources, including articulatory timing and shape. Recent research has shown that visual cortex encodes a categorical representation of articulatory features and that visual speech can benefit both acoustic and phonetic feature processing separately. The present study advances the hypothesis that the shape of the articulators specifically influences the categorization of auditory speech in terms of its phonetic features. We tested this by linearly modeling electroencephalographic responses to natural, continuous speech (in noise) in terms of the acoustic and articulatory features of the speech. We compared the performance of these models in conditions where the speech was accompanied by a natural video of the speaker with their mouth visible, and a video where their mouth was covered by a dynamic ellipse obscuring articulatory shape but preserving dynamics. The dynamic mask reduced comprehension, neural processing of phonetic features, the associated multisensory benefits, and indices of visual-only linguistic processing over occipital scalp. Our findings support substantial visual involvement in speech comprehension, derived largely from the shape of the articulators. They also corroborate several proposals involving audiovisual speech processing hierarchy and the nature of the information contained in visible speech. HighlightsO_LIVisual speech provides at least two forms of information to enhance acoustic speech processing: redundant temporal dynamics and complementary articulatory information C_LIO_LICovering the mouth with a dynamic mask preserves horizontal and vertical lip movement information, but largely removes articulatory detail C_LIO_LIVisual speech with a mask preserves some general multisensory benefits but removes visual linguistic information and its ability to enhance auditory processing at the level of phonetic features. C_LI
Adenis, V.; Bartholomew, R. A.; Lee, J.-I.; Jung, A.; Brown, M. C.; Fried, S. I.; Lee, D. J.; Arenberg, J. G.
Show abstract
Modern cochlear implants (CIs) use pulsatile stimulation to restore hearing for individuals with severe hearing loss. CIs provide robust speech recognition in quiet but poorly represent temporal fine structure (TFS), needed for challenging listening situations. Analog stimulation preserves the acoustic waveform and may better encode TFS, yet it has not been evaluated combined with modern current-focusing strategies. We compared neural responses in the inferior colliculus (IC) evoked by CI stimuli consisting of 100 pulses/s biphasic pulse trains and 100 cycles/s sinusoidal analog stimulation with monopolar, bipolar, and tripolar electrode configurations in urethane-anesthetized guinea pigs. Following cochlear implantation, multiunit activity was recorded from the tonotopic axis of the central nucleus of the IC using 16-channel silicon probes. Detection thresholds, spread of excitation, vector strength, sustained response percentage, and temporal response properties were quantified. Analog stimulation consistently evoked significantly lower activation thresholds than pulsatile stimulation while maintaining comparable or sometimes narrower spatial selectivity across stimulation modes. In contrast, analog stimulation generated lower vector strength, larger tonic response components, and a pronounced level-dependent polarity effect. At low stimulus levels, responses were dominated by the cathodic phase of the sinusoidal waveform, whereas increasing stimulus level responses were elicited by both phases, producing synchronization at twice the stimulus frequency. These findings demonstrate that stimulation waveform strongly influences temporal coding while having relatively little effect on the spatial distribution of neural activation. These results provide a physiological basis for reexamining analog stimulation as an alternative strategy for cochlear implant sound coding.
Rizzi, R.; Stirn, J. R.; Eisenhut, Z.; Bidelman, G. M.
Show abstract
Listeners discretize the speech signal by assigning sounds to phonetic categories, though there is variability in how individuals accomplish categorization. Having more consistent categorization of sounds may be advantageous for understanding speech-in-noise (SIN). Though, it is unclear how different levels of neural processing in the auditory system reflect these perceptual differences. We recorded brainstem frequency-following responses (FFRs) and cortical event-related potentials (ERPs) while listeners actively labeled vowels along an acoustic-phonetic continuum using a visual analog scale. We computed intertrial consistency of neural responses to index the stability of listeners' neural speech representations across stimulus presentations. We also assessed how faithfully midbrain and cortical responses represented stimulus acoustics using representational dissimilarity matrices (RDMs) computed across all token pairs. Neural RDMs were then compared with acoustic and phonetic category RDMs to assess whether FFRs and ERPs carried gradient vs. categorical information of the speech signal. We found greater behavioral consistency during phoneme labeling was correlated with improved SIN scores. Neurally, we found greater cortical or subcortical consistency predicted greater behavioral consistency. RDMs revealed subcortical responses retained more acoustic details, while cortical responses more closely reflected abstract phoneme categories. Our findings reveal important benefits of perceptual consistency to other domains of speech perception. We find perceptual consistency is driven by more consistent encoding of speech at either a cortical or subcortical level. More consistent sensory processing could provide a more stable readout of the speech signal to higher cortical brain areas which could confer advantages to later perceptual processes downstream.
Krasovskaya, S.; Coughlan, J. M.; Teng, S.
Show abstract
Some blind individuals use echolocation, a skill that allows them to better navigate their environment using echoes from self-generated mouth clicks reflected off surrounding surfaces. Echolocation involves a complex interplay of sensory accumulation, information processing, dynamic prediction, motor planning and execution in real-time. Computational modeling offers a valuable approach to understanding the cognitive and neural mechanisms underlying echolocation performance, in particular the temporal dynamics of the process. We present a computational model of human echolocation behavior based on a Kalman filter, where we treat the echolocator as an active sensor that maintains an internal belief about the target's location and continuously refines it via echo feedback. The model, based on observations of echolocation in blind human experts, simulates the use of mouth clicks and returning echoes to localize and orient toward a target under varying conditions. In the experiment, the target is placed at a random azimuth in the frontal plane. An echolocator aims a series of mouth clicks in various directions and infers the target azimuth using acoustic information received from the click echoes. The system integrates three major components: (1) a simulation of echoacoustic interaural time differences (ITD) to estimate the relative head-target angle; (2) a Kalman filter that processes these ITDs to iteratively update probabilistic beliefs about target location and associated uncertainty; and (3) a motor control system that modulates head movements with the current belief state. The Kalman filter serves as a representation of the internal state of the observer, where its beliefs drive the direction of head rotation, and its uncertainty estimates drive head velocity adjustments. Model performance demonstrates that simple predictive computational approaches can reproduce key aspects of echo-guided sensorimotor learning, providing a framework that may be leveraged to develop biologically plausible models, advance understanding of best practices, and potentially improve intervention strategies.
Yang, Y.; Li, X.; Li, M.; Zhang, Y.; Chen, M.; Fan, F.; Wang, K.; Du, H.
Show abstract
Brydes whales (Balaenoptera edeni edeni) are nationally protected in China, and the waters around Weizhou Island in the Beibu Gulf support one of the countrys few regularly observed coastal groups. However, acoustic data for this population remain limited, and the potential effects of local vessel noise are poorly described. We conducted 16 vessel-based surveys around Weizhou Island and adjacent waters in January 2024 using a low-disturbance sailboat platform and passive acoustic recorders, with concurrent visual observations where possible. We identified 734 low-frequency signals classified as putative Brydes whale vocalizations and quantified their temporal and spectral parameters. Call duration was significantly negatively correlated with maximum frequency and center frequency, but not with minimum frequency or bandwidth. Comparisons with published records indicate that the recorded signals are most similar to vocalizations previously reported from juvenile Brydes whales or mother-calf pairs, although individual source attribution could not be confirmed. Speedboat passage significantly increased root-mean-square sound pressure levels, and the dominant noise band overlapped the frequency range of the recorded Brydes whale signals, indicating potential for acoustic masking. These results expand the bioacoustic baseline for Brydes whales in Chinese coastal waters and provide evidence relevant to the management of vessel activity and whale-watching tourism around Weizhou Island.
Husta, C.; Seijdel, N.; Drijvers, L.
Show abstract
Face-to-face communication requires listeners to attend, integrate, and weigh multiple communicative signals, including auditory speech, mouth movements, and co-speech gestures. The contribution of these signals may depend on the reliability of auditory input and the informativeness of the available signals. We utilized rapid invisible frequency tagging (RIFT) with EEG to examine how participants attend to and integrate these different signals in clear and adverse listening conditions. Participants watched videos of an actress producing clear or noise-vocoded sentences. Auditory speech was amplitude-modulated at 58Hz, while the luminance of the gesture and mouth regions was frequency-tagged at 63Hz and 65Hz. Degraded speech elicited stronger responses at the auditory tagged frequency, suggesting increased attentional gain to the auditory signal when listening was challenging. In contrast, clear speech elicited stronger responses at the gesture tagged frequency and a stronger 2Hz intermodulation response (65-63Hz), reflecting enhanced nonlinear coupling between mouth movements and gestures. Finally, in degraded speech, the informativeness of mouth movement, but not gesture, was associated with intermodulation strength, suggesting that the informativeness of mouth movements plays a greater role in multisensory interaction when listening is challenging. Our findings demonstrate that both signal reliability and informativeness shape multisensory integration during spoken language comprehension.
Hovenkamp, P. D. L.; van Walraven, L.; Ollevier, A.; van Oevelen, D.; van der Stappen, A. F.
Show abstract
The advancement in deep learning techniques has made Convolutional Neural Networks (CNNs) a powerful tool for the fully automated classification of zooplankton images. In this study, we systematically investigate how network selection, colour information and differences in imaging instruments affect the classification of zooplankton images by comparing multiple state-of-the-art CNNs on images of zooplankton and marine snow from the in situ Continuous Particle Imaging and Classification Sensor (CPICS), Video Plankton Recorder (VPR), In Situ Ichtyoplankton Imaging System (ISIIS), and the on-board Plankton Imager (Pi-10). With differences between models of 7.8 to 19% in F1-score, we find that model selection strongly affects the classification performance, with EfficientNetV2S showing the most reliable overall performance. Moreover, differences between model architectures are largest for the least abundant classes (<100 labeled images), which implies that when these are present, careful model selection is most beneficial. The high image quality of the Pi-10 strongly increases the performance for the least abundant classes compared to the other instruments. In addition, we find a significant correlation (r = 0.597) between ImageNet the performance and F1-score on zooplankton images, which implies that more generally, a model that performs well on ImageNet will perform well for zooplankton classification. Colour information increases the F1-score of the best performing classifier with 2.8%, but provides a stronger benefit (25% F1-score) for classes with <100 images. The overall performance increase of colour information is less than expected and questions the advantage of recording colour information for zooplankton.
Letchumanan, J. S.; Gandhi, S.; Yin, H.; Blackman, S.; Fabbri, J.; Konofagou, E.; Kessler, D.; Shepard, K.
Show abstract
Point-of-care ultrasound has transformed bedside diagnostics, yet current systems remain limited by rigid form factors, bulky external electronics and the need for skilled operators. Here we report a conformable ultrasound imaging patch that integrates a 1024-channel CMOS ultrasound application-specific integrated circuit (ASIC) directly beneath a conformable piezocomposite transducer array. The 10 mm X 8 mm, 1024-element ASIC contains on-chip transmit and receive beamforming, reducing the effective off-chip channel count by 16X while preserving image fidelity. Fabricated on a flexible polyimide substrate and bonded using anisotropic conductive film, the patch operates untethered from conventional ultrasound consoles and requires only a laptop for control and data acquisition. The device supports focused, plane-wave and diverging-wave transmission with steering over {+/-}30{degrees} in azimuth and {+/-}15{degrees} in elevation, achieving peak-to-peak acoustic pressures up to 7 MPa at a 4.4-MHz center frequency (mechanical index of 1.7), within diagnostic safety limits. Phantom experiments demonstrate three-dimensional imaging with axial and lateral resolutions (in both XZ and YZ planes) of 0.5 mm and 2 mm, respectively, and accurate contrast reproduction in tissue-mimicking phantoms. Human studies further demonstrate three-dimensional (3D) visualization of the internal jugular vein and carotid artery, as well as rib-shadow-free imaging of pleural motion during respiration. This work establishes a scalable architecture for chronic, wearable ultrasound imaging and highlights the potential of CMOS-integrated, conformable ultrasound systems for continuous physiological monitoring and remote diagnostics.
Downie, I.; Szyszka, P.; Hall, N. J.; Edwards, T. L.
Show abstract
In turbulent environments, odorants from different sources arrive at different times, potentially providing cues for odor source segregation. In several invertebrate species, short differences in odorant onset enable freely moving animals to discriminate odorant mixtures. In vertebrates, however, studies of sensitivity to odorant onset asynchrony have been conducted under highly constrained sampling conditions, such as with odor delivery tightly coupled to respiration. In this study, we investigated whether domestic dogs could detect odorant onset asynchrony in odorant mixtures under conditions that preserve key features of natural odor sampling. Dogs performed a discrimination task in which odor stimuli were presented as ongoing pulse trains that began independently of animal behavior, avoiding artificial synchronization of odor delivery with sniff cycles. Dogs were trained to discriminate between mixtures of two odorants with synchronous onsets and mixtures with asynchronous onsets. Of the dogs trained, one was able to discriminate odorant onset asynchronies as short as 633 ms. Dogs also displayed sensitivity to auditory stimulus onset asynchrony, discriminating auditory asynchronies as short as 30 ms. These results provide the first demonstration of temporal sensitivity in canine olfaction and the first evidence that vertebrates can use odorant onset asynchrony under conditions that permit free odor sampling.
Qiu, C.; Li, D.; Huo, H.; Mishra, A.; Li, C.; Yin, K.; Wang, N.; Chen, J.; Yao, R.; Margolin, E. J.; Lipkin, M. E.; Zhong, P.; Ni, X.; Yao, J.
Show abstract
Urinary stone disease is a common urological condition with increasing incidence, particularly in developed countries. Laser lithotripsy (LL) has become a preferred minimally invasive treatment due to its high precision and low tissue damage. Recent studies suggest that cavitation plays a critical role in stone damage during LL, and three-dimensional passive cavitation mapping (3D-PCM) has emerged as a promising tool for detecting these events. However, clinical translation of 3D-PCM remains challenging due to limitations in imaging depth, field of view (FOV), and procedural compatibility. Here, we present a large-FOV dual-modality imaging system (3D-PCM and B-mode ultrasound) based on a large-aperture planar ultrasound array. Through array optimization and model-based reconstruction, our system achieves an expanded FOV of ~40*40mm^2 at a clinically relevant imaging depth of ~110mm, while maintaining high spatial resolution of ~0.6 mm laterally and ~0.4 mm axially. In vivo experiments in a porcine model demonstrate that the reconstructed cavitation distribution correlates well with stone damage. Our technology has the potential to provide real-time treatment feedback during LL without disrupting the standard workflow.
Hassan, M. W.; Crook, K.; Gi, Y. J.; Lee, J.; Hossain, M. M.
Show abstract
Objective: This study aims to develop and validate a quantitative, depth-resolved anisotropy imaging framework that extends ARFI-based focal degree-of-anisotropy (DoA) estimation into two-dimensional mapping by modeling the depth-dependent relationship between shear modulus ratio (SMR) and peak displacement ratio (PDR). Methods: We propose APRIL (Adaptive Polynomial Regression for anisotropy Imaging via ARFI-induced DispLacements), a framework for quantitative, depth-resolved DoA imaging that adaptively selects polynomial regression or shape-preserving spline interpolation based on excitation PSF asymmetry. Training data were generated using an LS-DYNA3D + Field II simulation pipeline in homogeneous transversely isotropic media (SMR 0.9-4.9). Testing included shifted SMRs under varied acoustic conditions and three heterogeneous inclusion configurations (anisotropic inclusion in isotropic background and vice versa). Experimental validation was performed in an in-vivo murine tumor model over the time, ex-vivo chicken breast, and tissue-mimicking gelatin phantoms, using a Verasonics system with an L11-5v transducer. Results: APRIL achieved depth-resolved SMR prediction errors below 9% over 10-30 mm, with highest accuracy in the focal region (MAE 2.3%, RMSE < 0.1) and stable performance across PSF transition zones. In heterogeneous phantoms, it reconstructed anisotropy maps with SSIM up to 86% and MPE below 7%, accurately delineating inclusion boundaries. Under acoustic parameter variations, mean absolute errors remained below 10%, demonstrating robustness to system and tissue heterogeneity. Conclusion: APRIL enables robust, two-dimensional anisotropy imaging beyond focal estimates. Significance: The method provides a physically grounded and generalizable framework for clinically viable anisotropy biomarkers in muscle, tendon, kidney, tumor and breast tissues.
Gibbons, A.; Parnell, A.; Donohue, I.; Ogasawara, M.; Ross, S. R. P.-J.
Show abstract
O_LIMonitoring and limiting the spread of invasive species on islands requires efficient detection and population estimation methods. However, elusive species can be difficult to monitor using traditional methods, making autonomous approaches such as camera trapping and acoustic monitoring increasingly valuable. C_LIO_LIOn the island of Okinawa, Japan, the small Indian mongoose ( Urva auropunctata) threatens many native species since its introduction in 1910. Listed among the worlds worst invasive species, effective monitoring of U. auropunctata in Okinawa is critical. The Okinawa Environmental Observation Network (OKEON) uses camera traps to detect U. auropunctata, but success depends on precise placement. Though OKEON also includes a high-resolution acoustic monitoring programme, no audio classification model currently exists for U. auropunctata. Developing such a model could improve substantially our capacity to detect and manage the species. C_LIO_LIUsing sparse U. auropunctata vocalisations collected from camera trap videos, we built a lightweight Convolutional Neural Network distilled from a more complex model for classifying contact calls and alarm calls of U. auropunctata. Our distilled model performed similarly to the full model at detecting vocalisations from training data, but was considerably faster. C_LIO_LIWe applied the distilled classifier to [~]486 hrs of audio collected over eight years from southern Okinawa, where we successfully detected U. auropunctata a handful of times in each year of recording. In spite of strong model performance on test data, our model did not transfer well to unseen data, perhaps owing to the rarity of U. auropunctata calls and consequent small training dataset size, limiting its utility for ecological monitoring. C_LIO_LIPractical implication. The use of sparse audio data from camera trap videos to train an acoustic classifier had limited utility to detect the rarely vocalising U. auropunctata from passive acoustic monitoring data. We provide several recommendations for enhancing classifier performance to provide robust actionable insights into the distribution and spread of U. auropunctata, and aid targeted conservation efforts for Okinawas threatened biodiversity. C_LI